Skip to content

GH-3696: Cache ParsedVersion in FileMetaData and use it in fromParquetMetadata - #3700

Open
asifsmohammed wants to merge 4 commits into
apache:masterfrom
asifsmohammed:gh-3696-cache-parsed-version
Open

GH-3696: Cache ParsedVersion in FileMetaData and use it in fromParquetMetadata#3700
asifsmohammed wants to merge 4 commits into
apache:masterfrom
asifsmohammed:gh-3696-cache-parsed-version

Conversation

@asifsmohammed

@asifsmohammed asifsmohammed commented Aug 1, 2026

Copy link
Copy Markdown

Rationale for this change

VersionParser.parse(createdBy) is called from 7 production sites, all parsing the same constant string from FileMetaData.getCreatedBy(). In fromParquetMetadata, this happens R×C times (once per column per row group) during footer metadata conversion. Since FileMetaData is
constructed once per file and already stores the createdBy string, it is the natural place to parse once and cache the result.

This PR caches the parsed version and migrates the first (and hottest) call site — ParquetMetadataConverter.fromParquetMetadata — to use the cache, eliminating redundant VersionParser.parse and SemanticVersion.parse calls from the R×C inner loop.

What changes are included in this PR?

  • Add getWriterVersion() to FileMetaData with lazy-init via an immutable WriterVersionResult holder (thread-safe, double-checked locking)
  • Add shouldIgnoreStatistics(ParsedVersion, PrimitiveTypeName) overload to CorruptStatistics that uses the cached SemanticVersion from ParsedVersion directly
  • Refactor ParquetMetadataConverter.fromParquetMetadata to construct FileMetaData before the row-group loop and use the cached ParsedVersion in buildColumnChunkMetaData
  • Falls back to the String-based path when getWriterVersion() throws VersionParseException to preserve exact logging behavior

Are these changes tested?

Yes.

  • FileMetaDataTest — 6 tests covering valid, null, empty, unparseable version strings, and caching
  • CorruptStatisticsTest.testParsedVersionOverload — covers all branches of the new ParsedVersion overload including null, non-parquet-mr, empty version, invalid semver, corrupt, and fixed versions
  • TestParquetMetadataConverter — 70 existing tests pass (validates the refactored fromParquetMetadata path)

Are there any user-facing changes?

No breaking changes. Adds new public methods:

  • FileMetaData.getWriterVersion() — returns cached ParsedVersion, throws VersionParseException for unparseable strings
  • CorruptStatistics.shouldIgnoreStatistics(ParsedVersion, PrimitiveTypeName) — for callers that already have a parsed version

@wgtmac wgtmac left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for fixing this! If this gets checked in, the PR that fixes malformed stats checking can be closed, right?

this.keyValueMetaData =
unmodifiableMap(Objects.requireNonNull(keyValueMetaData, "keyValueMetaData cannot be null"));
this.createdBy = createdBy;
this.writerVersion = parseVersion(createdBy);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Seems reasonable. Do we want to initialize this field lazily?

@wgtmac wgtmac left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for pushing this cleaner direction. I think caching the parsed writer version in FileMetaData is the right foundation, but this PR is not a complete fix yet. It currently adds the cache, but does not migrate the production call sites that still parse created_by repeatedly. Please update the hot/footer/page/reader/rewrite paths to consume the cached ParsedVersion, and cover that behavior in tests.

@@ -42,6 +44,7 @@
private final MessageType schema;
private final Map<String, String> keyValueMetaData;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This cached field is transient final. After Java deserialization it will stay null even when createdBy is valid. Recompute lazily in getWriterVersion(), or add serialization handling and tests.

* @return the parsed writer version, or {@code null} if {@code createdBy} is null, empty, or unparseable
*/
@JsonIgnore
public ParsedVersion getWriterVersion() {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This adds the cached value, but no production call site uses it yet. Please migrate the callers listed in the issue, especially footer stats, page stats, reader init, rewrite, and lazy encrypted metadata paths.

return VersionParser.parse(createdBy);
} catch (RuntimeException | VersionParser.VersionParseException e) {
return null;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Returning null loses the difference between missing createdBy and parse failure. CorruptStatistics currently logs different warnings for these cases. Preserve the parse failure state, or keep enough context for migrated callers to preserve existing behavior.

FileMetaData meta =
new FileMetaData(SCHEMA, Collections.emptyMap(), "parquet-mr version 1.12.0 (build abc123)");

assertThat(meta.getWriterVersion()).isNotNull();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These tests only cover the getter. They do not prove the repeated parsing problem is fixed. Please add coverage for at least one migrated production path using the cached ParsedVersion instead of reparsing createdBy.

@asifsmohammed
asifsmohammed force-pushed the gh-3696-cache-parsed-version branch from d7c4f6d to 2f01073 Compare August 2, 2026 16:59
…dant parsing

Parse the createdBy version string once during FileMetaData construction
and cache the result as a transient field. This avoids redundant
VersionParser.parse() calls at every downstream call site (R×C times
during footer decode alone).
@asifsmohammed
asifsmohammed force-pushed the gh-3696-cache-parsed-version branch from 2f01073 to bc23cf5 Compare August 2, 2026 17:08
- Change writerVersion to lazy computation on first getWriterVersion() call
- Fixes deserialization correctness (transient fields recompute from createdBy)
- Add writerVersionParsed flag to avoid retrying on parse failure
- Document contract for distinguishing missing vs. unparseable in javadoc
@asifsmohammed

Copy link
Copy Markdown
Author

This adds the cached value, but no production call site uses it yet. Please migrate the callers listed in the issue, especially footer stats, page stats, reader init, rewrite, and lazy encrypted metadata paths.

These tests only cover the getter. They do not prove the repeated parsing problem is fixed. Please add coverage for at least one migrated production path using the cached ParsedVersion instead of reparsing createdBy.

@wgtmac I'd prefer to keep this PR focused on the caching foundation and migrate callers in a follow-up. The migration touches multiple files which is a larger change that's easier to review separately. The follow-up will include tests proving the production path uses the cached version.
If needed I can add fixes for 1 caller using the ParsedVersion in this PR or I create a separate PR to fix all callers using ParsedVersion in parallel to this PR. Wdyt? I have no concerns on fixing all of them in this PR itself.

@wgtmac wgtmac left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the quick response!

I agree with keeping this PR small and moving the full migration to follow-ups. However, I do think it is worth migrating at least one production caller so this PR provides a concrete fix, not only an unused cache.

Please also narrow the title and description and remove “Closes #3696”, since the remaining paths still parse created_by repeatedly.

private final Map<String, String> keyValueMetaData;
private final String createdBy;
private transient volatile ParsedVersion writerVersion;
private transient volatile boolean writerVersionParsed;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please replace writerVersion and writerVersionParsed with one private immutable WriterVersionResult field. Use a null field only for “not initialized”. The result should represent valid, missing, or invalid and retain the original parse exception. Initialize it once in a synchronized block. Keep only getWriterVersion() public: return the parsed version, return null for missing createdBy, and rethrow the cached exception for invalid createdBy so callers can preserve the existing fallback without reparsing.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have a couple of callouts on this approach, happy to get your input on this @wgtmac

The shouldIgnoreStatistics(ParsedVersion, PrimitiveTypeName) overload achieves full behavioral parity
with the String-based overload for all real-world inputs. There are two theoretical edge cases where
the log output differs (return value is always identical):

  1. Empty version field (e.g., "parquet-mr version (build abc)"): The log message appends
    ParsedVersion.toString() instead of the raw createdBy string, since the ParsedVersion overload
    doesn't have access to the original string.

  2. Non-empty but invalid semver (e.g., "parquet-mr version xyz (build abc)"): The old code threw
    SemanticVersionParseException caught by the outer catch block, logging via warnParseErrorOnce
    with a stack trace. The new code logs via warnOnce without a stack trace, since ParsedVersion's
    constructor already caught and discarded the exception internally.

Neither case occurs in practice as no known parquet writer produces such strings. For truly unparseable
createdBy strings (where VersionParser.parse itself fails), fromParquetMetadata falls back to
the String-based path via the useWriterVersion flag, preserving exact logging parity including the
stack trace.

…zed init

- Replace writerVersion + writerVersionParsed with single WriterVersionResult
- Null field means not-yet-initialized, MISSING for null/empty createdBy
- Double-checked locking with synchronized for thread-safe one-time init
- Rethrow cached VersionParseException so callers preserve existing fallback
- Use Strings.isNullOrEmpty for consistency with CorruptStatistics
Add shouldIgnoreStatistics(ParsedVersion, PrimitiveTypeName) overload to
CorruptStatistics that uses the pre-parsed and cached SemanticVersion from
ParsedVersion, eliminating redundant VersionParser.parse and
SemanticVersion.parse calls in the R×C hot path.

Refactor ParquetMetadataConverter.fromParquetMetadata to construct the
hadoop FileMetaData before the row-group loop and extract the cached
ParsedVersion once via getWriterVersion(). The loop now uses the
ParsedVersion-based buildColumnChunkMetaData overload, avoiding per-column
re-parsing. Falls back to the String-based path when getWriterVersion()
throws VersionParseException to preserve exact logging parity.
@asifsmohammed asifsmohammed changed the title GH-3696: Cache ParsedVersion in FileMetaData to eliminate redundant parsing GH-3696: Cache ParsedVersion in FileMetaData and use it in fromParquetMetadata Aug 5, 2026
Comment on lines +108 to +113
if (!writerVersion.hasSemanticVersion()) {
warnOnce("Ignoring statistics because created_by could not be parsed (see PARQUET-251): " + writerVersion);
return true;
}

SemanticVersion semver = writerVersion.getSemanticVersion();

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ParsedVersion eagerly parses and caches the SemanticVersion in its constructor, so getSemanticVersion() avoids the redundant SemanticVersion.parse(version.version) that the String-based overload previously performed on every call. The left and right spikes in flame graph are for parsing SemanticVersion twice.

Image

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants